Papers with lexical metrics
GPT-4V Cannot Generate Radiology Reports Yet (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are becoming multimodal, and GPT-4 models are supposed to possess advanced skills across a wide range of domains, including high-stakes scenarios such as medicine. |
| Approach: | They perform a systematic evaluation of GPT-4 in generating radiology reports across three chest X-ray report benchmarks: MIMIC-CXR, CheXpert Plus, and IU X ray. |
| Outcome: | The proposed model fails in lexical and clinical efficacy metrics . the distributions of model-predicted labels remain constant regardless of groundtruth conditions on the image, suggesting that the model is not interpreting chest X-rays meaningfully. |
Multi-Layered Evaluation Using a Fusion of Metrics and LLMs as Judges in Open-Domain Question Answering (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for comparing machine-generated answers with reference are not perfect in terms of accuracy or cost. |
| Approach: | They propose to summarize long answers and use shortened versions to improve evaluation . they propose a multi-layered evaluation methodology that integrates different metrics tailored to various scenarios . |
| Outcome: | The proposed method outperforms existing evaluation methods but is more cost-effective than existing methods. |
StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Motivational interviewing (MI) is a directive, client-centered counseling approach for eliciting clients' motivation for behavioral change. |
| Approach: | They propose a multi-LLM agent framework for controllable MI dialogue generation . therapist and client agents generate MI-coded utterances guided by MI codes . |
| Outcome: | The proposed framework can generate fluent dialogues with minimal intervention time and a high level of evaluation. |